Papers by Eng Siong Chng
A Unified Speaker Adaptation Approach for ASR (2021.emnlp-main)
Copied to clipboard
| Challenge: | Adapting a model to target speakers requires a lot of compute and may cause catastrophic forgetting to the existing speakers. |
| Approach: | They propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. |
| Outcome: | The proposed model outperforms baseline models with 20.58% relative WER reduction and surpasses finetuning method by 2.54% on target speaker adaptation. |
Adapting BERT for Word Sense Disambiguation with Gloss Selection Objective and Example Sentences (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have used pre-trained language models for domain adaptation or transfer learning to improve natural language processing performance. |
| Approach: | They propose to fine-tune word sense disambiguation on sequence-pair ranking task and to use existing WordNet examples to augment the model. |
| Outcome: | The proposed model achieves state-of-the-art on the English all-words benchmark datasets. |
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition (2023.acl-long)
Copied to clipboard
| Challenge: | Existing efforts to improve robustness of audio-visual speech recognition with visual information focus on audio modality . current approaches introduce noise adaptation techniques to improve reliability of AVSR task . |
| Approach: | They propose a visual-invariant modality to strengthen robustness of audio-visual speech recognition (AVSR) it can adapt to any testing noises without dependence on noisy training data, a.k.a., unsupervised noise adaptation. |
| Outcome: | The proposed method outperforms existing state-of-the-arts on visual speech recognition task under various noisy and clean conditions. |
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)
Copied to clipboard
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal fusion methods ignore inter-modality relationship, treat each modality equally, suffer sensor noise, and thus reduce multimodal learning performance. |
| Approach: | They propose a multimodal contrastive method to explore more reliable multimodal representations under the weak supervision of unimodal predicting. |
| Outcome: | The proposed method outperforms current state-of-the-art multimodal learning methods on image-text classification benchmarks UPMC-Food-101 and N24News. |
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition (2023.acl-long)
Copied to clipboard
| Challenge: | Audio-visual speech recognition (AVSR) leverages multimodal signals to understand human speech. |
| Approach: | They propose an adversarial network to refine frame-level modality-invariant representations to bridge the distribution gap between modalities. |
| Outcome: | The proposed approach outperforms the state-of-the-art on public benchmarks LRS3 and LRS2 on the modalities of AVSR. |
InTriage: Intelligent Telephone Triage in Pre-Hospital Emergency Care (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing TT processes face challenges such as incomplete data collection, communication barriers, and manual errors, leading to high over-triage and under-triages rates. |
| Approach: | They propose to use an AI-driven multilingual TT system to provide decision support for triage. |
| Outcome: | The proposed system achieves word error rate of 14.57% for speech recognition and an F1 score of 73.34% for key information extraction. |
CASSI: Contextual and Semantic Structure-based Interpolation Augmentation for Low-Resource NER (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for text augmentation suffer from annotation corruption for token-level tasks like NER. |
| Approach: | They propose a novel augmentation scheme that generates high-quality contextually diverse augmentations while avoiding annotation corruption. |
| Outcome: | The proposed scheme outperforms existing methods at multiple low resource levels, in multiple languages, and for noisy and clean text. |